Papers with acoustic models

11 papers
CBAL: Context-Based Agentic Learning for Speaker Diarization Segmentation Refinement (2026.acl-srw)

Copied to clipboard

Challenge: Speaker diarization systems produce segmentation errors that degrade transcript readability and downstream applications.
Approach: They propose a framework that refines segmentation boundaries in diarized scripts . they use a lightweight LLM agent to reason about merge decisions .
Outcome: The proposed framework achieves 93.4% accuracy across 359 applied merges and reduces segment count by 6.1%.
Common Phone: A Multilingual Dataset for Robust Acoustic Modelling (2022.lrec-1)

Copied to clipboard

Challenge: Current state-of-the-art acoustic models can easily comprise more than 100 million parameters.
Approach: They propose to train a gender-balanced, multilingual corpus from 76.000 contributors via Mozilla’s Common Voice project to perform phonetic symbol recognition and validate the quality of the generated phonetic annotation.
Outcome: The proposed model can perform phonetic symbol recognition and validate the quality of the generated phonetic annotation.
Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Automatic pronunciation assessment is typically performed by acoustic models trained on audio-score pairs.
Approach: They propose a zero-shot, textual description-based Pronunciation Assessment approach that utilizes human-readable representations of speech signals fed into an LLM to assess pronunciation accuracy and fluency.
Outcome: The proposed approach is cost-efficient and competitive in performance . it significantly improves the performance of conventional audio-score-trained models on out-of-domain data .
Automatic Speech Recognition for Gascon and Languedocian Variants of Occitan (2024.lrec-main)

Copied to clipboard

Challenge: a new system for automatic speech recognition is being developed for two main Occitan dialects . the difficulty lies in the fact that Occitian is a less-resourced language .
Approach: They propose to develop an automatic speech recognition system for two Occitan dialects . they use Kaldi, acoustic models, and Whisper to create a model from corpora .
Outcome: The proposed system is based on Kaldi and Whisper for two main Occitan dialects . the system is more robust than previous systems, and the results are promising .
Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local Modeling (2023.emnlp-main)

Copied to clipboard

Challenge: Singing Voice Synthesis (SVS) synthesizes pleasing vocals based on music scores and lyrics . current acoustic models ignore the significance of local modeling within the sequence and the hard-to-synthesize parts in the predicted mel-spectrogram .
Approach: They propose a method to enhance local modeling in the acoustic model by focusing on phoneme tokens located before and after the phoneme.
Outcome: The proposed method improves local modeling in the acoustic model by focusing on the hard-to-synthesize parts of the predicted mel-spectrogram.
Quantifying Language Variation Acoustically with Few Resources (2022.naacl-main)

Copied to clipboard

Challenge: acoustic models represent linguistic information based on massive amounts of data.
Approach: They examine the model's ability to distinguish low-resource (Dutch) regional varieties by extracting embeddings from hidden layers and dynamic time warping.
Outcome: The proposed model outperforms transcription-based models without phonetic transcriptions on the basis of only six seconds of speech.
VoxCommunis: A Corpus for Cross-linguistic Phonetic Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Until recently, the movement towards large-scale cross-linguistic phonetic research has been limited.
Approach: They propose to use the VoxCommunis Corpus to facilitate cross-linguistic phonetic research . corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments .
Outcome: The VoxCommunis Corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . the corpus is free to download and use under a CC0 license .
Developing Resources for Automated Speech Processing of Quebec French (2020.lrec-1)

Copied to clipboard

Challenge: acoustic models for automatic segmentation of Quebec French are not available for all languages . linguistic resources are developed to perform phonetic annotations in Quebec French . physical characteristics of speech can be observed in the production of sounds .
Approach: They propose to use a French lexicon to train automatic QF segmentation models . they adapt existing pronunciation dictionary and acoustic model from existing ones .
Outcome: The proposed tools perform the full process of speech segmentation in Quebec French.
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)

Copied to clipboard

Challenge: Using acoustic data, we develop automatic speech recognition systems for three low resource languages.
Approach: They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework.
Outcome: The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut.
Improving Speech Recognition for the Elderly: A New Corpus of Elderly Japanese Speech and Investigation of Acoustic Modeling for Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: In an aging society, a highly accurate speech recognition system is needed for use in electronic devices for the elderly but this cannot be achieved using conventional speech recognition systems due to the unique features of the speech of elderly people.
Approach: They construct a new corpus of elderly Japanese speech from existing Japanese speech corpora and train them using existing data.
Outcome: The proposed models achieve word error rates (WER) as low as 13.38%, exceeding the results of the previous study.
SamróMur MilljóN: An ASR Corpus of One Million Verified Read Prompts in Icelandic (2024.lrec-main)

Copied to clipboard

Challenge: samrómur is a crowdsourcing web application designed to collect speech data for the advancement of language technologies in Icelandic.
Approach: They propose to use a crowdsourcing web application to collect and verify Icelandic speech data for automatic speech recognition (ASR) they introduce a dataset comprising one million audio clips from the application .
Outcome: The proposed system can produce high-quality speech data for Icelandic . the proposed system is based on a crowdsourced web application built on Mozilla's Common Voice .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations